Cell Genomics
○ Elsevier BV
All preprints, ranked by how well they match Cell Genomics's content profile, based on 172 papers previously published here. The average preprint has a 0.18% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Srivastava, J.; Ovcharenko, I.
Show abstract
A major challenge in deciphering the complex genetic landscape of Polycystic Ovary Syndrome (PCOS) lies in the limited understanding of how susceptibility loci drive molecular mechanisms across diverse phenotypes. To address this, we integrated molecular and epigenomic annotations from proposed causal cell-types and employed a deep learning (DL) framework to predict cell-type-specific regulatory effects of PCOS risk variants. Our analysis revealed that these variants affect key transcription factor (TF) binding sites, including NR4A1/2, NHLH2, FOXA1, and WT1, which regulate gonadotropin signaling, folliculogenesis, and steroidogenesis across brain and endocrine cell-types. The DL model, which showed strong concordance with reporter assay data, identified enhancer-disrupting activity in approximately 20% of risk variants. Notably, many of these variants disrupt TFs involved in androgen-mediated signaling, providing molecular insights into hyperandrogenemia in PCOS. Variants prioritized by the model were more pleiotropic and exerted stronger downregulatory effects on gene expression compared to other risk variants. Using the IRX3-FTO locus as a case study, we demonstrate how regulatory disruptions in tissues such as the fetal brain, pancreas, adipocytes, and endothelial cells may link obesity-associated mechanisms to PCOS pathogenesis via neuronal development, metabolic dysfunction, and impaired folliculogenesis. Collectively, our findings highlight the utility of integrating DL models with epigenomic data to uncover disease-relevant variants, reveal cross-tissue regulatory effects, and refine mechanistic understanding of PCOS.
Ahmed, O. Y.; Saravanan, N.; Rovsing, A. B.; Simpson, D.; Devarajan, A.; Gunn, S.; Singh, T.; Lappalainen, T.; Sanjana, N. E.
Show abstract
Over the past two decades, genome-wide association studies (GWAS) have identified thousands of trait- and disease-associated loci. However, the mechanistic understanding of these loci remains incomplete, which limits our ability to understand gene regulation and cellular programs underlying complex traits, predict disease risk, and develop therapeutics targeted to root causes. Here, we describe the current challenges for using GWAS to prioritize variants for functional follow-up experiments. These challenges span multiple domains, including limitations in data sharing and harmonization, limitations of statistical and functional fine-mapping, and the ambiguity in the added value of emerging deep learning frameworks for variant effect prediction as a complementary approach alongside traditional statistical genetics methods. We analyze these variant prioritization methods and suggest a multi-modal approach for resolving GWAS loci to a focused set of high-confidence variants for functional exploration. Fully realizing the potential of GWAS will require harmonized summary statistics and broader sharing of in-sample linkage disequilibrium (LD) data to enable robust and scalable causal variant prioritization.
Montero, J. J.; Trozzo, R.; Rad, R.
Show abstract
Despite their recognized role in biology, a majority of the [~]100,000 lncRNA genes remain functionally uncharacterized. In a recent study (Liang WW et al., Transcriptome-scale RNA-targeting CRISPR screens reveal essential lncRNAs in human cells, Cell, 2024), Liang et al. utilized the RNA nuclease Cas13d to perturb [~]6,200 lncRNAs in fitness screens across five cell lines - thereby identifying 778 lncRNAs with broad or context-specific essentiality. However, previous screens reported a lower proportion of essential lncRNAs. To investigate this discrepancy, we re-analysed Liang et al.s data and found that 68.1% of gRNAs causing fitness defects have off-targets in essential protein-coding genes. This caused numerous false-positive hits, particularly among lncRNAs classified as broadly essential. Off-target effects also compromise the studys validation efforts, including experiments combining single-cell transcriptomics and lncRNA-perturbations, which confirm the downregulation of off-target protein-coding genes identified in our analyses. The large number of false-positive hits reported by Liang et al. undermines the studys biological conclusions and endangers future research building on these data, if not considered.
Chen, H.-H.; Cai, Y.; Graff, M.; Zhu, W.; Petty, L. E.; Highland, H. M.; Roshani, R.; Polikowsky, H. G.; Lorenz, A. S.; Frankel, E.; Landman, J. M.; Anwar, M. Y.; Franson, E. E.; Haessler, J.; Avery, C. L.; Young, K. L.; Fernandez-Rhodes, L.; Reiner, A. P.; Peters, U.; Gordon-Larsen, P.; Gamazon, E. R.; Pereira, A. C.; Kooperberg, C.; Huff, C. D.; Fisher-Hoch, S. P.; McCormick, J. B.; North, K. E.; Below, J. E.
Show abstract
Genetic variants that influence transcript abundance, called expression quantitative trait loci (eQTL), are fundamental to understanding gene regulation and disease etiology. However, eQTL studies have overlooked the influence of the ancestral origin of a gene on its regulation. We implemented a new statistical framework that maps local ancestry-specific regulatory variation, revealing pervasive ancestry-specific effects on gene expression in Hispanic/Latino and African American populations. Enriched in open chromatin regions, these variants better explain genetic disease risk in these populations. The widespread heterogeneity of local ancestry-based eQTL effects offers mechanistic explanations for inconsistencies in genomic and multi-omic studies across populations. Our findings expand existing models of gene regulation and the importance of applying local genomic context in genetic studies to advance precision medicine and address health disparities.
Jones, A. G.; Connelly, G. G.; Dalapati, T.; Wang, L.; Schott, B.; SAN ROMAN, A. K.; Ko, D. C.
Show abstract
Humans display sexual dimorphism across many traits, but little is known about underlying genetic mechanisms and impacts on disease. We utilized single-cell RNA-seq of 480 lymphoblastoid cell lines to identify 1200 genes with significantly sex-biased expression. While reproducibility was highest among LCL datasets, 71% were found to be sex-biased in at least one GTEx tissue, with a core dataset of 21 genes displaying sex-biased expression across all datasets and tissues examined. While 7.7% of sex-biased genes can be directly explained by differences in the number of sex chromosomes, most sex-biased genes (79%) are targets of transcription factors that display sex-biased expression. FOSL1, ZNF730, ZFX, and ZNF726 appear to make the largest contribution to this based on machine learning and linear modeling approaches, and all four of these transcription factors are regulated by the number of X chromosomes. Further, by testing the difference in slopes of conditionally independent expression quantitative trait loci (eQTL) identified in each sex separately, we identified 2,390 sex-biased eQTL (sb-eQTL) across the genome. While evidence of replication in an independent dataset was modest, permutation analysis demonstrated that sb-eQTL identified using real sex was more likely to have concordant direction of effect. These sb-eQTL are enriched in over 100 GWAS phenotypes, including many loci associated with female-biased autoimmune diseases such as multiple sclerosis. Our results demonstrate widespread genetic impacts on sexual dimorphism and identify possible mechanisms and clinical targets for sex differences in diverse diseases.
Jia, Q.; Lam, M.; Wang, L.; Zhao, F.; Tang, H.; Sarashetti, P.; Li, Z.; Wong, E.; SG10K_Health Consortium, ; Tan, P.; Sim, X.; Ngeow, J.; Lee, J.; Cheng, C.-Y.; Chee, M. L.; Lim, W. K.; Chin, C. W. L.; Karnani, N.; Chong, Y. S.; Sim, W. C.; Lim, C. W.; Bertin, N.; Liu, J.
Show abstract
Tandem repeats (TRs) are implicated in over 70 Mendelian disorders and likely contribute to the "missing heritability" of complex traits and diseases, yet TR variations in Asian populations remain poorly characterized. Here, we constructed an Asian-specific SG10K-TR catalog by leveraging the SG10K_Health Dataset, comprising 916,274 autosomal TR loci genotyped in 9,490 individuals of Chinese (5,528), Malay (1,824), Indian (2,108), and other ancestries (30). Using a novel integrative measure for both repeat length and frequency variations, TRDDS, we found that population-level TR variations are selectively constrained in coding and promoter regions, whereas the enrichment of TRs with high population diversity was observed in regulatory sites with low chromatin accessibility and pathways related to neuronal functions. We also identified candidate TRs under selection that predominantly targets neuronal and synaptic architecture. Analysis of linkage disequilibrium (LD) patterns revealed that TRs are often poorly tagged by small variants, although we identified 123 candidate functional TRs that may underlie association signals previously attributed to nearby noncoding SNPs. Finally, TR-based GWAS of six anthropometric and lipid traits identified ten loci with genome-wide significant associations, including two novel loci for BMI (LINC02817) and height (UNC45B), and a TR variant as causal candidate for a known GWAS locus at HMGCR for LDL. Together, this study establishes a critical Asian-specific TR resource and highlights the fundamental role of TR diversity in driving evolutionary neuroplasticity and shaping the genetic architecture of complex traits.
Qu, H.-Q.; Ostberg, K.; Slater, D. J.; Wang, F.; Snyder, J.; Hou, C.; Connolly, J. J.; March, M.; Glessner, J. T.; Kao, C.; Hakonarson, H.
Show abstract
BackgroundType 1 diabetes (T1D) exhibits sex differences in genetic risk, yet most genetic studies treat sex as a covariate rather than a potential modifier of risk. We hypothesized that sex-stratified genome-wide association studies (GWAS) would uncover sex specific genetic architecture and improve risk prediction for T1D. MethodsWe performed GWAS in 6,599 T1D cases (3,483 males, 3,109 females, 7 undetermined) and 12,350 controls (6,665 males, 5,658 females, 27 undetermined) of European ancestry, testing both additive and additive-by-sex interaction models. We then conducted GWAS separately in males and females. For mechanistic insights into sex-specific effects, we generated single-cell RNA-sequencing (scRNA-seq) profiles of peripheral blood mononuclear cells (PBMCs) from nine matched male-female pediatric pairs of European ancestry. Finally, we tested male-, female-, and standard (all-samples) polygenic risk scores (PRS) in an independent cohort (471 T1D cases, 2,300 controls), and compared their performance by receiver operating characteristic (ROC) analysis. ResultsSex-stratified analyses identified 215 genome wide significant SNPs (P<5x10-8) exhibiting significant heterogeneity between sexes: 119 male-specific, 94 female-specific, and two shared SNPs at HLA-B (rs2249932 and rs2249934). Integration of scRNA-seq data pinpointed 41 genes with sex-specific T1D associations that also showed differential expression between males and females in particular cell types. In the independent cohort, sex specific PRS significantly outperformed the combined PRS: in males, AUC=0.668 versus 0.623 ({Delta}=0.045; DeLongs p<2.2x10-16); in females, AUC=0.719 versus 0.635 ({Delta}=0.084; DeLongs p<2.2x10-16). ConclusionsSex-stratified GWAS reveal novel T1D risk loci influenced by sex. Incorporating sex-specific effect sizes into PRS markedly enhances risk discrimination, underscoring the value of sex-aware genetic analyses for precise prediction and intervention in T1D.
Liu, S.; Zheng, H.; Gu, Y.; Yang, Z.; Liu, Y.; Wei, Y.; Guo, X.; Chen, Y.; Hu, L.; Chen, X.; Zhang, F.; Chen, G.-B.; Qiu, X.; Huang, S.; Zhen, J.; Wei, F.
Show abstract
The gestational period, spanning approximately 40 weeks from fertilization to birth, is fundamental to human reproduction. Health monitoring during this period involves systematic prenatal and postpartum examinations, guided by indicators collectively termed gestational phenotypes under the national medical insurance framework. Despite their clinical importance, the genetic basis of these phenotypes and their links to later-life health outcomes remain poorly understood. In this large-scale genetic study, we analyzed 122 gestational phenotypes in 121,579 Chinese pregnancies, encompassing anthropometric metrics, blood biomarkers, and common gestational complications and outcomes. We identified 3,845 genetic loci, including 1,893 novel loci, and uncovered gestation-specific genetic effects in 23 phenotypes, with proportions ranging from 0% to 100%. These loci were enriched in pathways related to hormonal regulation, cell growth and immune function. Longitudinal genome-wide association analyses of repeated measures across 24 complete blood cell phenotypes revealed significant gene-by-gestational timing interactions for 17.8% of loci across five gestational and postpartum periods. Mendelian Randomization of 220 mid- and late-life phenotypes identified 73 causal associations between gestational phenotypes and chronic diseases risks. These findings provide critical insights into the genetic architecture of human gestational phenotypes and their implications for long-term health, providing a foundation for advancing population health during gestation. Visualization of results is available at https://monn.pheweb.com/.
Gu, Y.; Zheng, H.; Wang, P.; Liu, Y.; Guo, X.; Wei, Y.; Yang, Z.; Cheng, S.; Chen, Y.; Hu, L.; Chen, X.; Zhang, Q.; Chen, G.; Wei, F.; Zhen, J.; Liu, S.
Show abstract
Gestational diabetes mellitus (GDM), a heritable metabolic disorder and the most common pregnancy-related condition, remains understudied regarding its genetic architecture and its potential for early prediction using genetic data. Here we conducted genome-wide association studies on 116,144 Chinese pregnancies, leveraging their non-invasive prenatal test (NIPT) sequencing data and detailed prenatal records. We identified 13 novel loci for GDM and 111 for five glycemic traits, with minor allele frequencies of 0.01-0.5 and absolute effect sizes of 0.03-0.62. Approximately 50% of these loci were specific to GDM and gestational glycemic levels, distinct from type 2 diabetes and general glycemic levels in East Asians. A machine learning model integrating polygenic risk scores (PRS) and prenatal records predicted GDM before 20 weeks of gestation, achieving an AUC of 0.729 and an accuracy of 0.835. Shapley values highlighted PRS as key contributors. This model offers a cost-effective strategy for early GDM prediction using clinical NIPT.
Reales, G.; Pullin, J. M.; Manipur, I.; Vigorito, E.; Wallace, C.
Show abstract
Colocalisation analysis is extensively applied across diverse GWAS and molecular QTL datasets to identify candidate causal genes. We systematically characterised large-scale colocalisation results across eQTL studies varying in cellular granularity and sample size, with the goal of providing design and interpretation recommendations. We found 34-50% of GWAS hits colocalised, and were more likely to colocalise if they were located nearer genes and had a more common lead variant. We also found over 50% of colocalisations were found in only one cell type. This led to an inherent trade-off: while high granularity studies tended to have smaller sample sizes and lower eQTL discovery, each eQTL from these high-granularity datasets were more likely to colocalise, reflecting cell-type specificity. On the other hand, lower granularity studies achieved larger sample size and higher eQTL discovery, leading to detection of the greatest total number of colocalisations, particularly for lower frequency GWAS lead variants. This suggests large, high granularity studies will be needed to identify remaining colocalisations. Of the peaks that colocalised, 37-47% did so with multiple genes, suggesting coregulation of the GWAS trait, horizontal pleiotropy, or false positives. However, sensitivity analyses indicated that even extremely stringent significance thresholds did not substantially reduce multi-gene colocalisations, arguing against widespread false discovery. Integration of enhancer-promoter interaction data provided evidence for coregulation among multi-colocalising eGenes. While disentangling causality from horizontal pleiotropy will ultimately require experimental perturbation, triangulation using different sources of observational data is likely to be necessary, provided careful consideration is taken to identify biases and missing data that may influence gene prioritisation.
Reales, G.; Wallace, C.
Show abstract
Genome-wide association studies (GWAS) have been a crucial tool in genomics and an example of applied reproducible science principles for almost two decades.1 Their output, summary statistics, are especially suited for sharing, which in turn enables new hypothesis testing and scientific discovery. However, GWAS summary statistics sharing rates have been historically low due to a lack of incentives and strong data sharing mandates, privacy concerns and standard guidelines.2 Albeit imperfect, citations are a key metric to evaluate the research impact. We hypothesised that data sharing might benefit authors through increased citation rates and investigated this using GWAS catalog3 data. We found that sharers get on average ~75% more citations, independently of journal of publication and impact factor, and that this effect is sustained over time. This work provides further incentivises authors to share their GWAS summary statistics in standard repositories, such as the GWAS catalog.
Harris, R. M.; Whitfield, T.; Blanton, L. V.; Skaletsky, H.; Blumen, K.; Hyland, P.; McDermott, E.; Summers, K.; Hughes, J. F.; Jackson, E.; Teglas, P.; Liu, B.; Chan, Y.-M.; Page, D. C.
Show abstract
The origins of sex differences in human disease are elusive, in part because of difficulties in separating the effects of sex hormones and sex chromosomes. To separate these variables, we examined gene expression in four groups of trans- or cisgender individuals: XX individuals treated with exogenous testosterone (n=21), XY treated with exogenous estradiol (n=13), untreated XX (n=20), and untreated XY (n=15). We performed single-cell RNA-sequencing of 358,426 peripheral blood mononuclear cells. Across the autosomes, 8 genes responded with a significant change in expression to testosterone, 34 to estradiol, and 32 to sex chromosome complement with no overlap between the groups. No sex-chromosomal genes responded significantly to testosterone or estradiol, but X-linked genes responded to sex chromosome complement in a remarkably stable manner across cell types. Through leveraging a four-state study design, we successfully separated the independent actions of testosterone, estradiol, and sex chromosome complement on genome-wide gene expression in humans.
Gonzalez Rivera, W. G.; Liu, Y.; Ma, N.; Kim, J.; D'Antonio, M.; Dube, U.; Frazer, K. A.; Gymrek, M.
Show abstract
Genetic studies have largely focused on homogeneous populations, limiting our understanding of the genetic architecture of complex traits in admixed individuals. The advent of diverse biobanks like the All of Us Research Program (AoU) and computationally efficient local ancestry inference (LAI) methods now enable admixture mapping (ADM) at biobank scale. Here, we used two orthogonal LAI methods (GNOMIX and FLARE) to characterize local ancestry in the entire AoU v7.1 cohort (n=230,019). We then used GNOMIX labels to identify associations between African (AFR) and Native American (NAT) local ancestry with 29 quantitative traits. We first analyzed All of Us v7.1 data across African (n=49,797) and Admixed American (n=40,327) cohorts, which identified 97 significant local ancestry associations (65 AFR, 32 NAT). These include strong known signals, such as an association between AFR ancestry at the DARC locus and white blood cell traits and between NAT ancestry at the BUD13/APOE5/ZPR1 locus and triglycerides, as well as additional signals not previously associated with local ancestry. We observed that trait associations with AFR local ancestry are largely consistent across the African and Admixed American cohorts, but that several AFR signals reach genome-wide significance exclusively in Admixed Americans. Grouping associations by trait category revealed distinct ancestral patterns: all endocrine, renal, and 75.0% of liver signals were driven by associations with NAT ancestry, whereas white blood cell (90.9%), red blood cell (65.1%), and lipid (66.7%) signals were largely associated with AFR ancestry, possibly reflecting different population-driven environmental exposures throughout history. Finally, we performed a second round of analysis, comprising the largest ADM study to date, on the entire AoU v7.1 cohort in which we pooled individuals from all ancestries. Despite evidence of confounding due to population structure, summary statistics for pooled results showed strong correlation (r>0.98) with those from single ancestry analysis and detected 2.7x fold more signals, most of which passed at least nominal significance and all of which showed consistent effect directions in the single ancestry cohorts. Overall, these results demonstrate the power of using large, admixed cohorts to gain new insights into the relationship between local ancestry and the genetic architecture of medically relevant complex traits.
Henry, A.; Senabouth, A.; Tyebally, R.; Bowen, B.; Allen, P. C.; Spenceley, E.; Sagi-Zsigmond, E.; Cuomo, A. S. E.; Fan, J.; Huang, H. L.; Tanudisastro, H. A.; Xue, A.; de Lange, K. M.; Figtree, G. A.; Hewitt, A. W.; MacArthur, D. G.; Powell, J. E.
Show abstract
Genome-wide association studies (GWAS) have been instrumental in uncovering the genetic basis of complex traits. When integrated with expression quantitative trait loci (eQTL) mapping, they can elucidate how risk loci influence traits through gene regulatory mechanisms. Recent single-cell eQTL (sc-eQTL) studies suggest that genetic effects on gene expression are often cell type- and subtype-specific, but such datasets have so far been underpowered for causal inference. Here, we leverage results from sc-eQTL mapping in the TenK10K project, comprising 154,932 common variant sc-eQTL across 28 immune cell types derived from matched whole-genome sequencing (WGS) and single-cell RNA-sequencing (scRNA-seq) of over 5 million peripheral blood mononuclear cells (PBMCs) from 1,925 individuals. We present a catalogue of cell type-specific causal effects of gene expression on 53 diseases (spanning 58,058 causal associations across 8,672 genes and 28 cell types), and 31 biomarker traits (spanning 674,764 causal associations across 16,078 genes and 28 cell types). By quantifying polygenic enrichment at both the single-cell and cell-type levels, we identify distinct immune cell contributions to both immune-related and systemic conditions. We demonstrate differential polygenic enrichment of Crohns disease and COVID-19 amongst dendritic cell subtypes, and high activity of B cell interferon II response in SLE. Integration with clinical drug development data reveals that therapeutic compounds targeting gene-trait associations identified in this study are three times more likely to have secured regulatory approval. Using Crohns disease as a motivating example, we demonstrate how population-based sc-eQTL data can pinpoint risk loci, effector genes and cell types, complementing findings from disease-focused tissue samples. Our findings provide a foundational resource for understanding the cell type-specific genetic architecture of disease and for guiding therapeutic discovery.
Blanton, L. V.; San Roman, A. K.; Wood, G.; Buscetta, A.; Banks, N.; Skaletsky, H.; Godfrey, A. K.; Pham, T. T.; Hughes, J. F.; Brown, L. G.; Kruszka, P.; Lin, A. E.; Kastner, D. L.; Muenke, M.; Page, D. C.
Show abstract
Recent in vitro studies of human sex chromosome aneuploidy showed that the Xi ("inactive" X) and Y chromosomes broadly modulate autosomal and Xa ("active" X) gene expression in two cell types. We tested these findings in vivo in two additional cell types. Using linear modeling in CD4+ T cells and monocytes from individuals with one to three X chromosomes and zero to two Y chromosomes, we identified 82 sex-chromosomal and 344 autosomal genes whose expression changed significantly with Xi and/or Y dosage in vivo. Changes in sex-chromosomal expression were remarkably constant in vivo and in vitro across all four cell types examined. In contrast, autosomal responses to Xi and/or Y dosage were largely cell-type-specific, with up to 2.6-fold more variation than sex-chromosomal responses. Targets of the X- and Y-encoded transcription factors ZFX and ZFY accounted for a significant fraction of these autosomal responses both in vivo and in vitro. We conclude that the human Xi and Y transcriptomes are surprisingly robust and stable across the four cell types examined, yet they modulate autosomal and Xa genes - and cell function - in a cell-type-specific fashion. These emerging principles offer a foundation for exploring the wide-ranging regulatory roles of the sex chromosomes across the human body.
Wright, S. N.; Yang, J.; Ideker, T.
Show abstract
While both common and rare variants contribute to the genetic etiology of complex traits, whether their impacts manifest through the same effector genes and molecular mechanisms is not well understood. Here, we systematically analyze common and rare variants associated with each of 373 phenotypic traits within a large biological knowledge network of gene and protein interactions. While common and rare variants implicate few shared genes, they converge on shared molecular networks for more than 75% of traits. We demonstrate that the strength of this convergence is influenced by core factors such as trait heritability, gene mutational constraints, and tissue specificity. Using neuropsychiatric traits as examples, we show that common and rare variants impact shared functions across multiple levels of biological organization. These findings underscore the importance of integrating variants across the frequency spectrum and establish a foundation for network-based investigations of the genetics of diverse human diseases and phenotypes.
Long, Y.; Ni, X.; Chen, T.; Hong, Q.; Wang, J.; Wang, C.; Huang, Z.; Xu, H.; Sun, M.; Pang, J.; Choi, J.; Zhang, T.; Long, E.
Show abstract
Even genetically identical cells in a homogeneous environment can exhibit heterogeneous mRNA abundance because of widely unavoidable random fluctuations, typically referred to as gene expression noise. Recent studies showed that noise, not just a nuisance, is indeed involved in cellular activities (e.g., immune response), evolutionary processes, and diseases mechanisms. However, determinants of the gene expression noise and its functional role in variations of human complex traits remain largely unexplored. Here, we established an atlas of gene expression noise from 1.23 million human peripheral blood cells of 981 individuals, identifying its age- and gender-dependent pattern. We then identified 10,770 independent expression noise quantitative trait loci (enQTLs) for 6,743 unique enGenes (genetically driven gene expression noise) across 7 immune cell types. Most enQTLs were distinct from expression quantitative trait loci (eQTLs) and showed differential enrichment of functional elements across the genome. Colocalization of enQTLs with trait-associated genetic loci interpreted previously unexplained loci and pinpointed novel putative genes underlying hematopoietic traits and autoimmune diseases. Overall, this study unravels the genetic determinants of gene expression noise and implicates as a previously underappreciated mechanism underlying variation of human complex traits and diseases.
Castel, S. E.; Tluway, F.; Emde, A.-K.; Smyth, N.; Karim, M.; Sengupta, D.; Gray, O. O.; Hendershott, M.; LeBaron von Baeyer, S.; Burke, E.; Kaewert, S.; Nguyen, K.; Choma, S.; Mashaba, G.; Micklesfield, L.; Kabudula, C.; Kahn, K.; Gomez-Olive, F.; Tollman, S.; Choudhury, A.; Mpangase, P.; Hazelhurst, S.; Wasik, K. A.; Yerges-Armstrong, L.; Ramsay, M.
Show abstract
Functional genomics resources are critical for interpreting human genetic studies, however they are predominantly from European-ancestry individuals. Here we present the Southern African Blood Regulatory (SABR) resource, a map of blood regulatory variation that includes three South Eastern Bantu-speaking groups. Using paired whole genome and blood transcriptome data from over 600 individuals, we map the genetic architecture of 40 blood cell traits derived from deconvolution analysis, as well as expression, splice, and cell type interaction quantitative trait loci. We comprehensively compare SABR to the Genotype Expression (GTEx) Project and characterize the thousands of African-enriched and African-specific regulatory variants mapped. Finally, we demonstrate the increased utility of SABR for interpreting African association studies by identifying putatively causal genes and molecular mechanisms through colocalization analysis of 83 blood-relevant traits from the PAN-UK Biobank. Importantly, we make full SABR summary statistics publicly available to support the African genomics community.
Oelen, R.; de Vries, D. H.; Brugge, H.; Gordon, G.; Vochteloo, M.; Ye, C. J.; Westra, H.-J.; Franke, L.; van der Wijst, M. G. P.
Show abstract
Gene expression and its regulation can be context-dependent. To dissect this, using samples from 120 individuals, we single-cell RNA-sequenced 1.3M peripheral blood mononuclear cells exposed to three different pathogens at two time points or left unexposed. This revealed thousands of cell type-specific expression changes (eQTLs) and pathogen-induced expression changes (response QTLs) that are influenced by genetic variation. In monocytes, the strongest responder to pathogen stimulations, genetics also affected co-expression of 71.4% of these eQTL genes. For example, the pathogen recognition receptor CLEC12A showed many such co-expression interactions, but only in monocytes after 3h pathogen stimulation. Further analysis linked this to interferon-regulating transcription factors, a finding that we recapitulated in an independent cohort of patients with systemic lupus erythematosus, a condition characterized by increased interferon activity. Altogether, this study highlights the importance of context for gaining a better understanding of the mechanisms of gene regulation in health and disease.
Trang, K. B.; Sharma, P.; Cook, L.; Mount, Z.; Thomas, R. M.; Kulkarni, N. N.; Pahl, M. C.; Pippin, J. A.; Su, C.; Kaestner, K. H.; O'Brien, J. M.; Wagley, Y.; Hankenson, K. D.; Jermusyk, A.; Hoskins, J. W.; Amundadottir, L. T.; Xu, M.; Brown, K. M.; Anderson, S. A.; Yang, W.; Titchenell, P. M.; Seale, P.; Zemel, B. S.; Chesi, A.; Romberg, N.; Levings, M. K.; Grant, S. F.; Wells, A. D.
Show abstract
Genome-wide association studies (GWAS) have identified genetic links to autoimmune disorders, but lack detail on causal elements. We generated 3D genomic datasets of promoter-focused Capture-C, Hi-C, ATAC-seq, and RNA-seq across 57 human cell types integrated with GWAS of 16 autoimmune traits. These data allowed us to map disease-associated variants to their effector genes and identify impacted cell types more effectively than using 1D genomic features or eQTL approaches. Most variants implicated by 3D cis-regulatory architectures are trait-specific, while half the target genes are shared across multiple disorders and cell types, leading to enrichment of similar biological networks. This indicates complex genetic diversity converges on shared targets, yet unique pathways were identified offering avenues for targeted therapies. We pharmacologically validated squalene synthase, a cholesterol biosynthetic enzyme encoded by the FDFT1 gene implicated by our approach and eQTL in multiple sclerosis and systemic lupus erythematosus, as a novel immunomodulatory drug target controlling T cell inflammatory cytokine production and aiding B cell antibody production in a human lymphoid organoid model. These data offer a comprehensive resource for understanding gene cis-regulatory mechanisms, and the analyses shed light on how autoimmune-associated variants regulate gene expression, function, and pathology across diverse tissues and cell types.